Papers with audio-language models

8 papers
SoundMind: RL-Incentivized Logic Reasoning for Audio-Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Recent large language models have demonstrated impressive reasoning abilities, but their extension to the audio modality remains underexplored.
Approach: They propose a rule-based reinforcement learning algorithm to equip LALMs with robust reasoning capabilities.
Outcome: The proposed algorithm improves on the SoundMind benchmark.
AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for audio-centric interaction have impeded advancements in this field . AIR-Bench evaluates LALMs' ability to understand audio signals and interact with humans .
Approach: They propose a benchmark to evaluate the ability of large audio-language models to understand audio signals . they use 19 tasks with approximately 19k single-choice questions to examine single-task ability .
Outcome: The proposed framework evaluates the ability of large audio-language models to understand audio signals and interact with humans in the textual format.
Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for music question answering do not systematically evaluate reasoning across tracks.
Approach: They propose a dataset and benchmark for multi-track comparative question answering . they construct 36,519 comparative QA items over 12,173 track pairs .
Outcome: The proposed dataset and benchmark for multi-track comparative question answering is based on the Jamendo-QA dataset.
Unlocking Large Audio-Language Models for Interactive Language Learning (2026.findings-eacl)

Copied to clipboard

Challenge: Computer-Assisted Pronunciation Training (CAPT) systems provide unintuitive feedback that lacks actionable guidance.
Approach: They propose to use audio-language models to provide more user-friendly feedback for pronunciation training.
Outcome: The proposed model outperforms baselines on mispronunciation detection and suggestion generation.
FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio Pretraining (2026.acl-long)

Copied to clipboard

Challenge: Existing audio-language models excel at clip-level understanding but struggle with frame-level tasks.
Approach: They propose a novel training paradigm that advances both clip- and frame-level alignment in CLAP with heterogeneous data.
Outcome: The proposed training paradigm improves both clip- and frame-level alignment in CLAP with heterogeneous data.
Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models (2026.acl-long)

Copied to clipboard

Challenge: Emotion is a central dimension of spoken communication, yet we lack a mechanistic account of how LALMs encode it internally.
Approach: They propose to use emotion-sensitive neurons in large audio-language models to study their interpretations.
Outcome: The proposed models show that they can be used to make decisions on emotion . the results show that the ESNs exhibit non-uniform clustering with partial cross-dataset transfer .
EMO-RL: Emotion-Rule-Based Reinforcement Learning Enhanced Audio-Language Model for Generalized Speech Emotion Recognition (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in reinforcement learning (RL) have shown promise in improving LALMs’ reasoning abilities, but their performance in affective computing tasks remains suboptimal.
Approach: They propose a framework incorporating reinforcement learning with two key innovations: Emotion Similarity-Weighted Reward (ESWR) and Explicit Structured Reasoning (ESR).
Outcome: The proposed framework improves LALMs' reasoning abilities on MELD and IEMOCAP datasets and shows strong generalization.
iKnow-audio: Integrating Knowledge Graphs with Audio-Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Contrastive language-audio pretraining models learn by aligning audio and text in a shared embedding space.
Approach: They propose a framework that integrates knowledge graphs with audio-language models to provide robust semantic grounding.
Outcome: iKnow-audio improves disambiguation of acoustically similar sounds and reduces prompt engineering.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations